Papers with Chinese social media

9 papers
The UIR Uncertainty Corpus for Chinese: Annotating Chinese Microblog Corpus for Uncertainty Identification from Social Media (L18-1)

Copied to clipboard

Challenge: Uncertainty identification is an important semantic processing task, critical to the quality of information in terms of factuality in many NLP techniques and applications.
Approach: They propose to annotate Chinese microblogs with an open uncertainty corpus . they propose to use contextual uncertain semantics rather than traditional cue-phrases to identify uncertainty .
Outcome: The proposed corpus can be used to identify uncertainty in social media texts.
Emotion Cause Extraction on Social Media without Human Annotation (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies have focused on extracting emotion causes from news articles, but lack of fine-grained annotations has limited the ECE task.
Approach: They propose a new ECE framework that extracts emotion causes from social media data without relying on human annotations.
Outcome: The proposed framework achieves high extraction performance and generalizability without relying on human annotations.
Named Entity Recognition for Social Media Texts with Semantic Augmentation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches for named entity recognition suffer from data sparsity problems when conducted on short and informal texts.
Approach: They propose a neural-based approach to named entity recognition for social media texts . they obtain augmented semantic information from a large-scale corpus and encode it .
Outcome: The proposed approach outperforms existing approaches on three social media datasets.
Autoencoder as Assistant Supervisor: Improving Text Representation for Chinese Social Media Text Summarization (P18-2)

Copied to clipboard

Challenge: Existing abstractive text summarization models learn a semantic representation of the source text and the summaries from it.
Approach: They evaluate the model on a popular Chinese social media dataset and compare it to other models.
Outcome: The proposed model achieves state-of-the-art performance on a popular Chinese social media dataset.
Topic-Aware Neural Keyphrase Generation for Social Media Language (P19-1)

Copied to clipboard

Challenge: Existing methods to extract words from source posts to form keyphrases do not exploit latent topics.
Approach: They propose a sequence-to-sequence-based neural keyphrase generation framework . it allows absent keyphrases to be created, and it allows joint modeling of latent topic representations .
Outcome: The proposed model outperforms extraction and generation models without exploiting latent topics.
Chinese MentalBERT: Domain-Adaptive Pre-training on Social Media for Chinese Mental Health Text Analysis (2024.findings-acl)

Copied to clipboard

Challenge: Existing models for language analysis are inadequate for specialized domains like psychology.
Approach: They have enriched a Chinese social media database with psychological lexicons to enhance its applicability to psychological text analysis.
Outcome: The proposed model performed better on six public datasets and provided relevant predictions given the masked sentences.
Can Large Language Models Understand Internet Buzzwords Through User-Generated Content (2025.acl-long)

Copied to clipboard

Challenge: Existing methods for generating definitions of internet buzzwords rely on user-generated content, such as posts and reviews, to understand them.
Approach: They propose a method to generate accurate buzzword definitions using UGC as examples.
Outcome: The proposed method mirrors human language learning skills and can produce more accurate definitions.
MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Social media is a key platform for emotional expression, yet deep learning lacks flexibility and interpretability.
Approach: They propose to use Chinese social media to train interpretable mental health instruction datasets to test models' ability to explain their decisions.
Outcome: The proposed models outperform deep learning and LLMs on three mental health downstream tasks and demonstrate their potential for clinical applications.
Image Matters: A New Dataset and Empirical Study for Multimodal Hyperbole Detection (2024.lrec-main)

Copied to clipboard

Challenge: linguistic detection of hyperbole is an important part of understanding human expression . studies on hyperbolic expressions focus on text modality, but social media can be used to detect it .
Approach: They propose to use a multimodal detection dataset to study hyperbole detection . they treat text and image as two modalities and evaluate pre-trained encoders .
Outcome: The proposed dataset is constructed from five different keywords and shows its performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations